DPO: Direct Preference Optimization
DPO is an optimization that simplifies the two-stage pipeline containing reward model training in RLHF and PPO post-training into a single supervised loss function.
DPO is an optimization that simplifies the two-stage pipeline containing reward model training in RLHF and PPO post-training into a single supervised loss function.
Date: 2026-09-14 Mon